You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
I have a model that I am trying to train where the loss does not go down. I have a custom image set that I am using. These images are 106 x 106 px (black and white) and I have two (2) classes, Bargraph or Gels. These two classes are very different. I have run the Cifar10 dataset and it did reduce the loss, but I am very confused as to why my model will always predict only one class for everything.
Xtrain is a numpy array of images (which are numpy arrays), Ytrain is a numpy array of arrays ([0,1] or [1,0]) the shapes look like this:
Right now I am just doing very small training sets (I tried doing 1000 examples as well, with similar results).
I have also tried RMS and SDG with large and small learning rates.
What else can I try ?
Check that you are up-to-date with the master branch of Keras. You can update with:
pip install git+git://github.com/fchollet/keras.git --upgrade --no-deps
If running on Theano, check that you are up-to-date with the master branch of Theano. You can update with:
pip install git+git://github.com/Theano/Theano.git --upgrade --no-deps
beyhangl, feay1234, vgovindarajulu, fedecaccia, RGaonkar, will-rice, toobaimt, BastiQ, r-boutin, Lxrd-AJ, and 24 more reacted with thumbs up emojiiFe1er, peterwashington, and CShorten reacted with laugh emojihadifar, ongss, zoink, peterwashington, and CShorten reacted with hooray emojihadifar, BastiQ, titusaj, zoink, awandzel, hafiz031, peterwashington, CShorten, and katejarne reacted with heart emojiAll reactions
Tried it, it stopped after 2 epochs. Here are the results
The first class is from 0 to 999 and the second class is from 1000 to 1999
I tried to predict right at the border and got all [1,0]. Shuffling the training set should not matter should it?
After reading some blogs, looks like the batch size is important, because if our data is not shuffled it will learn one class for a few batches and then another class for a few batches. Similarly My loss seems to stay the same, here is an interesting read on the loss function. I really am still unsure as to what I may be doing wrong.
Here are a few things I tried:
number of layers (reduction)
size of the filters (reduction)
SGD learning rate from 0.000000001 to 0.1
SGD decay to 1e-2
Batch size
Different images
Shuffling the images around
I am really unsure as to what I can do to get my loss to go down. Any other ideas?
srikar2097, ciozhang, tzrm, ghvn7777, ChenluJi, ArturoDeza, momonala, zccoder, jackvial, PeterPanUnderhill, and 14 more reacted with thumbs up emojiptran1203 reacted with laugh emojiptran1203 reacted with heart emojiAll reactions
Miss activation (e.g. relu) after Convolution2D. I use your network on cifar10 data, loss does not decrease but increase. With activation, it can learn something basic.
Network is too shallow. It's hard to learn with only a convolutional layer and a fully connected layer. Try Alexnet or VGG style to build your network or read examples (cifar10, mnist) in Keras.
I recommend you to take some online courses about deep learning, it would be helpful.
Yevgnen, FoxerLee, momonala, Mike201456, ddvogt, Coldmaple, LIexcalibur, kwanCCC, artificialskills, padipadou, and 184 more reacted with thumbs down emojiLetiP, vmvargas, AliBaheri, ychennay, titusaj, vitorsantos95, cheezbuggah, chrizzzzy, katherineyun, amapic, and 6 more reacted with confused emojic7huang reacted with rocket emojithomasryck, vitorsantos95, SeyedHosseinTafakh, zxh3, awandzel, madpeh, and someshdev reacted with eyes emojiAll reactions
I am unsure as to what you mean in 1 for miss activation. Are you saying if you remove the activation the loss increases and when you use activation it learns?
For 2: I actually tried with a deeper network, but I figured since it was giving me no improvement, it may be best to simplify the model and troubleshoot with that.
I am wondering if this could be an issue with my data.
I can try to increasing the depth. Anything else I should look at?
is unnecessary because we do not need to shuffle the input (This was just a test to try and figure out why My network would not converge).
I still have problems with RMSprop. It quickly gains loss, and the accuracy goes to 0 (which to me is funky). I tried a few different SGDs and the one in my latest post seemed to work the best for me.
@kevkid I also meet your problem. I collect 1505 numbers pics as my dataset and use a simple model.
The valid_acc doesn't change. Do you give me some advices. Thanks.
@kevkid
Have you found the solution now?I have met the same problem to yours.
Have you tried to change the ' momentum=1.9 '.I found that this problem may connected to the argument named 'momentum' in SGD optimizter.
I did't find the solution yet but when I changed the momentum to 0.5, the loss changed.But after several epoch , the loss did not change again......
hope this can help you !
Hello, I used vgg19 architecture to classify my data set into 2 classes but the problem is the value of accuracy doesn't change after 30 iteration and I don't what is the problem and this my code:
data,Label = shuffle(immatrix,label, random_state=2)
train_data = [data,Label]
print (train_data[0].shape) # the train data
print (train_data[1].shape) # the labels of these data
#batch_size to train
batch_size = 16
My modest experience tells me that if you have only two classes use a dict in class_weight.
If you have more you'll get the error class_weight not supported for +3 dim.
A way to overcome this consists in adding sample_weight in fit() using a 2D weight array (one weight per timestep per sample), and adding sample_weight_mode="temporal" in compile()
It's not an elegant solution but it works. I'll be glad if someone has another answer.
Hey, i am having a similar problem i am trying to train a network to learn word embeddings using skip grams. i have a vocabulary of 256 and a sequence of about 166000 words. But when i train, the accuracy stays the same at around 0.1327 no matter what i do, i tried changing learning rates and batch_size. But no luck. This has happened every time i used keras. But it usually starts learning after tweaking the batch size a bit. But this one just doesn't work.
Here is the model:
def make_model(self,vocab_size=256,vec_dim=100):
model=Sequential()
model.add(Dense(vec_dim,activation="sigmoid",input_dim=vocab_size))
model.add(Dense(vocab_size,activation="sigmoid"))
sgd=SGD(lr=1.0)
model.compile(loss="categorical_crossentropy",optimizer=sgd,metrics=["accuracy"])
return model
def _get_callbacks(self):
earlystop=EarlyStopping(monitor="val_loss",min_delta=0.0001,patience=10,verbose=2)
checkpoint=ModelCheckpoint("checkpt.hdf5",period=10,verbose=2)
reducelr=ReduceLROnPlateau(monitor="val_loss",factor=0.1,patience=5,verbose=2)
return [earlystop,checkpoint,reducelr]
def train(self,model,X,y):
model.fit(X,y,nb_epoch=1000,callbacks=self._get_callbacks(),validation_split=0.1,verbose=2,batch_size=300)
model.save("model.hdf5")
X is a one hot vector of len 256 for every word
y in a one hot vector of len 256 representing the skip word in the context of X.
so for instance is the sequence is [2,6,5,7,9]
X will be
[5,5,5,5,7,7,7...]
y will be
[2,6,7,9,6,5,9...]
and so on for every word in the sequence.
I ve waited for a about 50 epochs and the acc still does not change.
Any idea what i am doing wrong? Ive faced this problem everytime ive used keras even when training other models like language modelling using RNNs text generation using LSTMs.
Hi guys, I am having a similar problem. I am training an LSTM model for text classification and my loss does not improve on subsequent epochs. I tried many optimizers with different learning rates. But same problem.
max_length = 275
X_train = sequence.pad_sequences(train_data_new, maxlen=max_length, padding='post')
X_test = sequence.pad_sequences(test_data_new, maxlen=max_length, padding='post')
y_train = []
y_test = []
# preparing y_test and y_train
for label in train_label:
if label == 'first':
y_train.append([1,0])
else:
y_train.append([0,1])
y_train = np.array(y_train)
for label in test_label:
if label == 'second':
y_test.append([1,0])
else:
y_test.append([0,1])
y_test = np.array(y_test)
# Create the model
rmsprop = RMSprop(lr=0.1)
sgd = SGD(lr=0.1)
model = Sequential()
model.add(Embedding(len(word_index) + 1, EMBEDDING_DIM,
weights=[embedding_matrix],
input_length=max_length,
trainable=True))
# model.add(Dropout(0.2))
model.add(LSTM(128, return_sequences=True))
# model.add(Dropout(0.2))
model.add(LSTM(64))
model.add(Dense(2, activation='sigmoid'))
model.compile(loss='categorical_crossentropy', optimizer=rmsprop, metrics=['accuracy'])
print(model.summary())
print('Training model...')
model.fit(X_train, y_train, validation_data=(X_test, y_test), epochs=10, batch_size=64)
# Final evaluation of the model
scores = model.evaluate(X_test, y_test, verbose=0)
print("Accuracy: %.2f%%" % (scores[1]*100))
Initially, it was default. Then I read on a similar issue page on stackoverflow where it told to alter learning rates. So I was trying to see the change on different values. Using lr=0.1 the loss starts from 0.83 and becomes constant at 0.69. When I was using default value, loss was stuck same at 0.69
Okay. I created a simplified version of what you have implemented, and it does seem to work (loss decreases). Here is the code you can cut and paste. Note that the first section is setting up the environment for reproducible results (which I provide at the end in my case). In your case, you may want to check a few things:
Is your input data making sense? It could be that the preprocessing steps (the padding) are creating input sequences that cannot be separated (perhaps you are getting a lot of zeros or something of that sort).
You might want to simplify your architecture to include just a single LSTM layer (like I did) just until you convince yourself that the model is actually learning something.
I hope this helps. Thanks.
Start: Set up environment for reproduction of results
import numpy as np
import tensorflow as tf
import random as rn
import os
os.environ['PYTHONHASHSEED'] = '0'
np.random.seed(42)
rn.seed(12345)
#single thread
session_conf = tf.ConfigProto(
intra_op_parallelism_threads=1,
inter_op_parallelism_threads=1)
from keras import backend as K
tf.set_random_seed(1234)
sess = tf.Session(graph=tf.get_default_graph(), config=session_conf)
K.set_session(sess)
End: Set up environment for reproduction of results
from keras.layers import LSTM, Dense, Embedding
from keras.models import Sequential
from keras.preprocessing import sequence
I encountered the problem while I was trying to finetune a pretrained VGGFace model, using keras_vggface.utils.preprocess_input as my custom preprocessing function.
def preprocess_input(x, data_format=None, version=1):
if data_format is None:
data_format = K.image_data_format()
assert data_format in {'channels_last', 'channels_first'}
if version == 1:
if data_format == 'channels_first':
x = x[:, ::-1, ...]
x[:, 0, :, :] -= 93.5940
x[:, 1, :, :] -= 104.7624
x[:, 2, :, :] -= 129.1863
else:
x = x[..., ::-1]
x[..., 0] -= 93.5940
x[..., 1] -= 104.7624
x[..., 2] -= 129.1863
elif version == 2:
if data_format == 'channels_first':
x = x[:, ::-1, ...]
x[:, 0, :, :] -= 91.4953
x[:, 1, :, :] -= 103.8827
x[:, 2, :, :] -= 131.0912
else:
x = x[..., ::-1]
x[..., 0] -= 91.4953
x[..., 1] -= 103.8827
x[..., 2] -= 131.0912
else:
raise NotImplementedError
return x
The problem seems to come from the scaling. I used preprocessing_function=keras_vggface.utils.preprocess_input and got into that problem. However, when I rescale it with 1/255. the problem is fixed. I think it may be that the pretrained model was trained additionally with a scaling factor to normalize it to [0,1], but the preprocessing function only gives us the mean so we know which means to subtract to center the data. I'd recommend you check if your scaling makes sense; a bad scaling of inputs into a Neural Network may cause your updates to either move very slowly (i.e. the derivative of the sigmoid function beyond -3 and +3 are near 0 and so your gradients are almost 0), or if you're using something like the ReLU function, the updates may be big (the derivative is 1) and a wrong update makes you jump pass the local minima very easily.
ALSO, if you're rescaling in python 2, make sure you have that dot in 1/255., or else all your inputs will be multiplied by 0 and you aren't making any updates!!!
In my case, It is the normalization problem:
(x_train, y_train), (x_test, y_test) = cifar10.load_data()
x_train = x_train.astype('float32')
x_test = x_test.astype('float32')
I am trying to Create DNN but it is not converging, any idea
model = Sequential()
model.add(Dense(5000, input_dim=5, activation='relu', kernel_regularizer=regularizers.l2(0.1)))
model.add(Dropout(0.1))
model.add(Dense(2000, kernel_regularizer=regularizers.l2(0.1),activation='relu'))
model.add(Dropout(0.1))
model.add(Dense(400, kernel_regularizer=regularizers.l2(0.1),activation='relu'))
model.add(Dropout(0.1))
model.add(Dense(1,activation='relu'))
I had a similar issue today when training on Google cloud GPU. I tried changing network architecture, weights, etc. The solution was to reset the TF graph: tf_reset_default_graph()
Somehow the GPU seemed to have a "memory" across different runs and was stuck at a local minima.
I had a similar issue today when training on Google cloud GPU. I tried changing network architecture, weights, etc. The solution was to reset the TF graph: tf_reset_default_graph()
Somehow the GPU seemed to have a "memory" across different runs and was stuck at a local minima.
Thx,I will have a try,hope it would work!
For me, theese 3 things did the trick:
Lower the learning rate (0.1 converges too fast and already after the first epoch, there is no change anymore). Just for test purposes try a very low value like lr=0.00001.
Check the input for proper value range and normalize it
Add BatchNormalization (model.add(BatchNormalization())) after each layer
SriHarsha-Paladugula, chnzhangrui, kangaroo02, aql315, amapic, keremakinli, and soans1994 reacted with thumbs up emojikeremakinli and KenzSem reacted with hooray emojiunnir reacted with confused emojikeremakinli, cagrisayir, KenzSem, and soans1994 reacted with heart emojikeremakinli reacted with rocket emojikeremakinli reacted with eyes emojiAll reactions
For me, this works:
Add BatchNormalization (model.add(BatchNormalization())) after each layer
Thanks @Fellfalla
Based on my own experience as a starter, one possible reason or bug in your model is that you probably used a wrong activation function, i.e. the way you activated your result at the last output layer, for example, if you are trying to solve a multi class proplem, usually we use softmat rather sigmoid, while sigmoid is meant to activate the output for binary task. And in this case, it's a binary application, therefore just change your activation function as sigmoid, you should not find such exception.
I had a model that did not train at all. It just stucks at random chance of particular result with no loss improvement during training. Loss was constant 4.000 and accuracy 0.142 on 7 target values dataset.
It become true that I was doing regression with ReLU last activation layer, which is obviously wrong.
Before I was knowing that this is wrong, I did add Batch Normalisation layer after every learnable layer, and that helps. However, training become somehow erratic so accuracy during training could easily drop from 40% down to 9% on validation set. Accuracy on training dataset was always okay.
Then I realized that it is enough to put Batch Normalisation before that last ReLU activation layer only, to keep improving loss/accuracy during training. That probably did fix wrong activation method.
However, when I did replace ReLU with Linear activation (for regression), no Batch Normalisation was needed any more and model started to train significantly better.
My data had 3 classes but last layer was Dense(1, activation='sigmoid') changing it to Dense(3, activation='sigmoid') made the loss change. With the 1 output neuron it didn't return any errors, just had a constant loss.
Also if you are training binary classifier, you can just use Dense(1, activation='sigmoid') as output with binary_crossentropy, instead of Dense(2, activation='sigmoid') with categorical_crossentropy